docs(eval): add README for running evals locally#249
Conversation
|
/fs-fix rebase and solve the conflicts |
|
🤖 Finished Fix · ✅ Success · Started 9:53 AM UTC · Completed 9:58 AM UTC |
Document prerequisites, required env vars, run commands, test case structure, and known issues for the eval framework. Covers both review and triage agents. Closes #247 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
🔧 Fix agent — iteration 1 (human-triggered)Rebased onto origin/main and resolved add/add conflict in eval/README.md by merging both versions into a cohesive document. Pre-existing test failure (prerequisites-creates-allowed-issue) is unrelated to this docs change. Fixed (1):
Tests: passed Decision points
Updated by fullsend fix agent |
93edc38 to
15d57ad
Compare
|
/fs-review |
|
🤖 Finished Review · ✅ Success · Started 10:06 AM UTC · Completed 10:14 AM UTC |
|
Looks good to me |
|
🤖 Finished Retro · ✅ Success · Started 10:20 AM UTC · Completed 10:32 AM UTC |
Retro: PR #249 — docs(eval): add README for running evals locallyTimeline
Total cost: ~$5.24 across 4 agent runs. What went well
What could improve
Autonomy evidenceThe review agent fully covered the human review on this docs-only PR — both approved with zero findings, zero delta. The correctness sub-agent's verification was thorough (cross-referencing env vars, CLI commands, and file paths against 6+ shell scripts). This adds evidence to the pattern tracked in fullsend-ai/fullsend#5289 (review agent autonomy on docs-only PRs). Related tracking issues: fullsend#4765, fullsend#5289. Existing issue evidence
Proposals filed |
Summary
eval/README.mddocumenting how to run agent evals locallyCloses #247
Test plan
eval/README.mdexists and renders correctlyeval.yamlconfigs are documentedeval/run-functional.shusage🤖 Generated with Claude Code
Closes #247
Post-script verification
agent/247-eval-readme)def8cb529ebf40b98eb0c0b67a2793b46f7648e4..HEAD)